Compute Chips

Silicon Photonics (CPO) Penetration Acceleration Curve in AI Cluster Optical Interconnects

As AI clusters scale to tens of thousands of accelerators and beyond, the limitations of traditional electrical interconnects become increasingly apparent. Bandwidth demands, reach requirements, and power constraints push data center architects to seek new ways to move bits efficiently between nodes. Silicon photonics, and in particular co‑packaged optics (CPO), has emerged as a critical technology for addressing these challenges.

The Cost-Benefit of GPU Cluster Migration from InfiniBand to Ethernet (RoCE)

As AI training clusters scale from dozens to thousands of GPUs, the interconnect fabric becomes one of the largest and most strategic line items in the infrastructure budget. For years, InfiniBand has been the de facto choice for high‑performance GPU clusters, particularly in large‑scale deep learning and HPC environments. At the same time, Ethernet with RDMA over Converged Ethernet (RoCE) has quietly matured, closing much of the performance gap while retaining the economic and operational advantages of mainstream Ethernet ecosystems.

Impact of Upgraded Reliability and Lifetime Test Standards on AI Chip Yields

As AI workloads move from experimental deployments to mission‑critical services, expectations around the reliability and lifetime of AI chips have risen sharply. Data centers, automotive systems, industrial controllers, and consumer devices now run AI models continuously, often under harsh thermal and electrical conditions. In response, chip makers and system integrators are upgrading reliability and lifetime test standards to ensure their devices can withstand years of heavy duty.

China’s AI Chip Localization: The Qualitative Leap from "Usable" to "Good"

Over the past decade, China’s push to localize its semiconductor stack has moved from an aspirational policy goal to an operational reality, especially in the realm of AI chips. Early domestic accelerators were often labeled “good enough” or merely “useful” – serviceable for certain workloads, but rarely the first choice for cutting‑edge model training or large‑scale deployment. Today, the conversation is shifting.

ASICs Eating into GPU Share: Broadcom and Marvell’s Golden Era

For much of the last decade, GPUs have been the default answer to almost any question about high‑performance compute and AI acceleration. They offered flexible parallelism, strong software ecosystems, and a simple story: one architecture, many workloads. That narrative is starting to fragment. In more data‑center racks and custom systems, application‑specific integrated circuits (ASICs) are quietly claiming sockets that might once have gone to GPUs.

AI Inference Chip Landscape: Startups Challenging Nvidia’s Triton Ecosystem

AI inference has quietly become the backbone of modern digital experiences: search, recommendation, content ranking, copilots, and vision systems all depend on running trained models efficiently and at scale. In this realm, Nvidia’s Triton Inference Server and its GPU‑centric ecosystem have established a powerful beachhead, offering a unified way to deploy and manage models across data‑center GPUs. Yet a growing wave of startups is challenging this dominance, not necessarily by duplicating Triton, but by reshaping hardware and software assumptions around inference.

The Profit Erosion Effect of Soaring Tape-Out Costs for China’s AI Chip Designers

In the last few years, China’s AI chip industry has moved from early experimentation to large‑scale commercial deployment, with dozens of design houses racing to tape out increasingly complex accelerators for data centers, edge devices, and vertical applications. At the same time, the cost of moving a design from layout to silicon—tape out—has risen sharply, driven by advanced process nodes, more complex packaging, and growing verification demands.

  • 1
  • 2
  • 3